NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

iSeqSearch: incremental protein search for iBlast/iMMSeqs2/iDiamond

https://doi.org/10.7717/peerj.19171

Yoo, Hyunwoo; Refahi, Mohammadsaleh; Polikar, Robi; Sokhansanj, Bahrad A; Brown, James R; Rosen, Gail L (April 2025, PeerJ)

BackgroundThe advancement of sequencing technology has led to a rapid increase in the amount of DNA and protein sequence data; consequently, the size of genomic and proteomic databases is constantly growing. As a result, database searches need to be continually updated to account for the new data being added. However, continually re-searching the entire existing dataset wastes resources. Incremental database search can address this problem. MethodsOne recently introduced incremental search method is iBlast, which wraps the BLAST sequence search method with an algorithm to reuse previously processed data and thereby increase search efficiency. The iBlast wrapper, however, must be generalized to support better performing DNA/protein sequence search methods that have been developed, namely MMseqs2 and Diamond. To address this need, we propose iSeqsSearch, which extends iBlast by incorporating support for MMseqs2 (iMMseqs2) and Diamond (iDiamond), thereby providing a more generalized and broadly effective incremental search framework. Moreover, the previously published iBlast wrapper has to be revised to be more robust and usable by the general community. ResultsiMMseqs2 and iDiamond, which apply the incremental approach, perform nearly identical to MMseqs2 and Diamond. Notably, when comparing ranking comparison methods such as the Pearson correlation, we observe a high concordance of over 0.9, indicating similar results. Moreover, in some cases, our incremental approach, iSeqsSearch, which extends the iBlast merge function to iMMseqs2 and iDiamond, provides more hits compared to the conventional MMseqs2 and Diamond methods. ConclusionThe incremental approach using iMMseqs2 and iDiamond demonstrates efficiency in terms of reusing previously processed data while maintaining high accuracy and concordance in search results. This method can reduce resource waste in continually growing genomic and proteomic database searches. The sample codes and data are available at GitHub and Zenodo (https://github.com/EESI/Incremental-Protein-Search; DOI:10.5281/zenodo.14675319).
more » « less
Free, publicly-accessible full text available April 28, 2026
Normalized Compression Distance for DNA Classification

https://doi.org/10.1145/3698587.3701490

Hearne, Gavin LA; Refahi, Mohammad S; Duan, Haozhe Neil; Brown, James R; Rosen, Gail L (November 2024, ACM)

Full Text Available
The Role and Applications of Artificial Intelligence in the Treatment of Chronic Pain

https://doi.org/10.1007/s11916-024-01264-0

Meier, Tiffany A; Refahi, Mohammad S; Hearne, Gavin; Restifo, Daniele S; Munoz-Acuna, Ricardo; Rosen, Gail L; Woloszynek, Stephen (August 2024, Current Pain and Headache Reports)

Full Text Available
Fragment databases from screened ligands for drug discovery (FDSL-DD)

https://doi.org/10.1016/j.jmgm.2023.108669

Wilson, Jerica; Sokhansanj, Bahrad A.; Chong, Wei Chuen; Chandraghatgi, Rohan; Rosen, Gail L.; Ji, Hai-Feng (March 2024, Journal of Molecular Graphics and Modelling)

Full Text Available
Interpretable and Predictive Deep Neural Network Modeling of the SARS-CoV-2 Spike Protein Sequence to Predict COVID-19 Disease Severity

https://doi.org/10.3390/biology11121786

Sokhansanj, Bahrad A.; Zhao, Zhengqiao; Rosen, Gail L. (December 2022, Biology)

Through the COVID-19 pandemic, SARS-CoV-2 has gained and lost multiple mutations in novel or unexpected combinations. Predicting how complex mutations affect COVID-19 disease severity is critical in planning public health responses as the virus continues to evolve. This paper presents a novel computational framework to complement conventional lineage classification and applies it to predict the severe disease potential of viral genetic variation. The transformer-based neural network model architecture has additional layers that provide sample embeddings and sequence-wide attention for interpretation and visualization. First, training a model to predict SARS-CoV-2 taxonomy validates the architecture’s interpretability. Second, an interpretable predictive model of disease severity is trained on spike protein sequence and patient metadata from GISAID. Confounding effects of changing patient demographics, increasing vaccination rates, and improving treatment over time are addressed by including demographics and case date as independent input to the neural network model. The resulting model can be interpreted to identify potentially significant virus mutations and proves to be a robust predctive tool. Although trained on sequence data obtained entirely before the availability of empirical data for Omicron, the model can predict the Omicron’s reduced risk of severe disease, in accord with epidemiological and experimental data.
more » « less
Full Text Available
Complet+: a computationally scalable method to improve completeness of large-scale protein sequence clustering

https://doi.org/10.7717/peerj.14779

Nguyen, Rachel; Sokhansanj, Bahrad A.; Polikar, Robi; Rosen, Gail L. (January 2023, PeerJ)

A major challenge for clustering algorithms is to balance the trade-off between homogeneity, i.e. , the degree to which an individual cluster includes only related sequences, and completeness, the degree to which related sequences are broken up into multiple clusters. Most algorithms are conservative in grouping sequences with other sequences. Remote homologs may fail to be clustered together and instead form unnecessarily distinct clusters. The resulting clusters have high homogeneity but completeness that is too low. We propose Complet+, a computationally scalable post-processing method to increase the completeness of clusters without an undue cost in homogeneity. Complet+ proves to effectively merge closely-related clusters of protein that have verified structural relationships in the SCOPe classification scheme, improving the completeness of clustering results at little cost to homogeneity. Applying Complet+ to clusters obtained using MMseqs2’s clusterupdate achieves an increased V-measure of 0.09 and 0.05 at the SCOPe superfamily and family levels, respectively. Complet+ also creates more biologically representative clusters, as shown by a substantial increase in Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) metrics when comparing predicted clusters to biological classifications. Complet+ similarly improves clustering metrics when applied to other methods, such as CD-HIT and linclust. Finally, we show that Complet+ runtime scales linearly with respect to the number of clusters being post-processed on a COG dataset of over 3 million sequences. Code and supplementary information is available on Github: https://github.com/EESI/Complet-Plus .
more » « less
Full Text Available
Predicting COVID-19 disease severity from SARS-CoV-2 spike protein sequence by mixed effects machine learning

https://doi.org/10.1016/j.compbiomed.2022.105969

Sokhansanj, Bahrad A.; Rosen, Gail L. (October 2022, Computers in Biology and Medicine)

Full Text Available
How Scalable Are Clade-Specific Marker K-Mer Based Hash Methods for Metagenomic Taxonomic Classification?

https://doi.org/10.3389/frsip.2022.842513

Gray, Melissa; Zhao, Zhengqiao; Rosen, Gail L. (July 2022, Frontiers in Signal Processing)

Efficiently and accurately identifying which microbes are present in a biological sample is important to medicine and biology. For example, in medicine, microbe identification allows doctors to better diagnose diseases. Two questions are essential to metagenomic analysis (the analysis of a random sampling of DNA in a patient/environment sample): How to accurately identify the microbes in samples and how to efficiently update the taxonomic classifier as new microbe genomes are sequenced and added to the reference database. To investigate how classifiers change as they train on more knowledge, we made sub-databases composed of genomes that existed in past years that served as “snapshots in time” (1999–2020) of the NCBI reference genome database. We evaluated two classification methods, Kraken 2 and CLARK with these snapshots using a real, experimental metagenomic sample from a human gut. This allowed us to measure how much of a real sample could confidently classify using these methods and as the database grows. Despite not knowing the ground truth, we could measure the concordance between methods and between years of the database within each method using a Bray-Curtis distance. In addition, we also recorded the training times of the classifiers for each snapshot. For all data for Kraken 2, we observed that as more genomes were added, more microbes from the sample were classified. CLARK had a similar trend, but in the final year, this trend reversed with the microbial variation and less unique k-mers. Also, both classifiers, while having different ways of training, generally are linear in time - but Kraken 2 has a significantly lower slope in scaling to more data.
more » « less
Full Text Available
Mapping Data to Deep Understanding: Making the Most of the Deluge of SARS-CoV-2 Genome Sequences

https://doi.org/10.1128/msystems.00035-22

Sokhansanj, Bahrad A.; Rosen, Gail L. (April 2022, mSystems)
Gaglia, Marta M. (Ed.)
ABSTRACT Next-generation sequencing has been essential to the global response to the COVID-19 pandemic. As of January 2022, nearly 7 million severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) sequences are available to researchers in public databases. Sequence databases are an abundant resource from which to extract biologically relevant and clinically actionable information. As the pandemic has gone on, SARS-CoV-2 has rapidly evolved, involving complex genomic changes that challenge current approaches to classifying SARS-CoV-2 variants. Deep sequence learning could be a potentially powerful way to build complex sequence-to-phenotype models. Unfortunately, while they can be predictive, deep learning typically produces “black box” models that cannot directly provide biological and clinical insight. Researchers should therefore consider implementing emerging methods for visualizing and interpreting deep sequence models. Finally, researchers should address important data limitations, including (i) global sequencing disparities, (ii) insufficient sequence metadata, and (iii) screening artifacts due to poor sequence quality control.
more » « less
Full Text Available
Correction: Candidate variants in DNA replication and repair genes in early-onset renal cell carcinoma patients referred for germline testing

https://doi.org/10.1186/s12864-023-09486-z

Demidova, Elena V.; Serebriiskii, Ilya G.; Vlasenkova, Ramilia; Kelow, Simon; Andrake, Mark D.; Hartman, Tiffiney R.; Kent, Tatiana; Virtucio, James; Rosen, Gail L.; Pomerantz, Richard T.; et al (December 2023, BMC Genomics)

Full Text Available

« Prev Next »

Search for: All records